Papers with automatic evaluation metrics

97 papers
HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability.
Approach: They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations.
Outcome: The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement .
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
GREEN: Generative Radiology Report Evaluation and Error Notation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability.
Approach: They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors.
Outcome: The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods.
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.
FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation (2023.tacl-1)

Copied to clipboard

Challenge: a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation is presented . FRMT is a type of style-targeted translation that uses labeled training data to perform tasks.
Approach: They propose a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation.
Outcome: The proposed model is based on two translations from English into Portuguese and Mandarin Chinese.
Evaluating Dialogue Generation Systems via Response Selection (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic evaluation metrics for open-domain dialogue systems correlate poorly with human evaluation.
Approach: They propose to construct response selection test sets with well-chosen false candidates to evaluate response generation systems via response selection.
Outcome: The proposed method correlates with human evaluation better than widely used metrics such as BLEU.
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages? (2024.acl-short)

Copied to clipboard

Challenge: a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources.
Approach: They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages.
Outcome: The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi.
Collaborative Document Simplification Using Multi-Agent Systems (2025.coling-main)

Copied to clipboard

Challenge: Document simplification requires complex factors such as technical terminology, metaphors, and overall coherence.
Approach: They propose a multi-agent framework for document simplification based on large language models that emulates the collaborative process of a human expert team through the roles played by multiple agents.
Outcome: The proposed framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification.
MathBuddy: A Multimodal System for Affective Math Tutoring (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing LLM-based conversational systems do not take into account the student’s affective states.
Approach: They propose an emotionally aware LLM-powered math tutor that models student emotions and maps them to relevant pedagogical strategies.
Outcome: The proposed model improves student engagement and learning effectiveness by 23 points using win rate and 3 points at an overall level using DAMR scores.
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation (2020.acl-main)

Copied to clipboard

Challenge: Standard language generation metrics have been shown to be ineffective for dialog evaluation.
Approach: They propose an unsupervised evaluation metric for dialog that trains unsupervised models to measure several desirable qualities of dialog.
Outcome: The proposed evaluation metric strongly correlates with human judgment on Topical-Chat and PersonaChat.
A Statistical Analysis of Summarization Evaluation Metrics Using Resampling Methods (2021.tacl-1)

Copied to clipboard

Challenge: Existing methods for summarization evaluations that approximate human judgments are lacking for accuracy and reliability.
Approach: They propose methods for calculating confidence intervals and running hypothesis tests for correlations using bootstrapping and permutation.
Outcome: The proposed methods show that the confidence intervals are wide, demonstrating high uncertainty in the reliability of automatic metrics.
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist (2023.acl-long)

Copied to clipboard

Challenge: a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human .
Approach: They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation .
Outcome: The proposed framework provides access to the evaluation tools for three NLG tasks.
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way (2022.emnlp-main)

Copied to clipboard

Challenge: Existing summarization datasets often have issues that seriously limit their usability.
Approach: They propose a faster but more straightforward approach to developing summarization benchmark data . they use a protocol that hires highly-qualified contractors to read stories and write original summaries from scratch .
Outcome: The proposed protocol is faster but more straightforward than scraping summaries from everyday text.
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model (2025.findings-acl)

Copied to clipboard

Challenge: Recent research has demonstrated that large language models (LLMs) can translate cultural elements in languages such as idioms and proverbs.
Approach: They propose to use large language models to translate culturally rooted proverbs in conversation and between languages with similar cultural backgrounds to compare their results.
Outcome: The proposed models can achieve good translation between languages with similar cultural backgrounds and outperform NMT models in proverb translation.
Location-Aware Visual Question Generation with Lightweight Models (2023.emnlp-main)

Copied to clipboard

Challenge: a novel task aims to generate engaging questions from location-aware information . a lightweight model can be used to generate such questions .
Approach: They propose a task to generate engaging questions from location-aware data . they represent location-based information with surrounding images and a GPS coordinate .
Outcome: The proposed method outperforms baselines regarding human evaluation and evaluation metrics.
Generating Reasonable and Diversified Story Ending Using Sequence to Sequence Model with Adversarial Training (C18-1)

Copied to clipboard

Challenge: Story generation is a challenging problem in artificial intelligence (AI) . previous work focused on learning statistical models of event sequences from large-scale text corpora .
Approach: They propose to use adversarial training to generate reasonable story endings . their model includes a generator that defines the policy of generating a story ending .
Outcome: The proposed model achieves better performance on the task of Story Cloze Test with an accuracy of 62.6% compared with state-of-the-art baseline methods.
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? (2025.acl-short)

Copied to clipboard

Challenge: Existing automatic evaluation metrics are based on procedures that diverge from human evaluation.
Approach: They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system.
Outcome: The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics.
Commonsense and Named Entity Aware Knowledge Grounded Dialogue Generation (2022.naacl-main)

Copied to clipboard

Challenge: Empirical results show that our proposed model outperforms the state-of-the-art methods in terms of both automatic evaluation metrics and human judgment.
Approach: They propose a model which uses large-scale commonsense and named entity based knowledge to ground dialogue on external knowledge and topic-specific knowledge associated with each utterance.
Outcome: The proposed model outperforms the state-of-the-art methods on two benchmark datasets.
On the Evaluation of Vision-and-Language Navigation Instructions (2021.eacl-main)

Copied to clipboard

Challenge: Existing instruction generators have not been evaluated using human wayfinders . BLEU, ROUGE, METEOR and CIDEr are ineffective for evaluating grounded navigation instructions.
Approach: They propose an instruction-trajectory compatibility model that operates without reference instructions to improve wayfinding performance.
Outcome: The proposed model shows the highest correlation with human wayfinding outcomes when scoring individual instructions.
Learning from Missing Relations: Contrastive Learning with Commonsense Knowledge Graphs for Commonsense Inference (2022.findings-acl)

Copied to clipboard

Challenge: Existing approaches to commonsense inference lack coverage and expressive diversity of commonsensense knowledge graphs.
Approach: They propose a framework that contrasts sets of semantically similar and dissimilar events . they propose 'solar' framework that can be used to learn commonsense inference .
Outcome: The proposed framework outperforms the state-of-the-art commonsense transformer on commonsensense inference by 1.84% on average among 8 metrics.
Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic dialog evaluation metrics are mostly reference-based . Existing models that measure self-reported user ratings are biased and variance among different users.
Approach: They propose an automatic evaluation model that automatically cleans self-reported user ratings as it trains on them.
Outcome: The proposed model achieves 89.2% accuracy in the dialog comparison task.
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that doctors can save significant amounts of time when using automatic note generation.
Approach: They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics.
Outcome: The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets.
Varifocal Question Generation for Fact-checking (2022.emnlp-main)

Copied to clipboard

Challenge: Recent question generation approaches assume that the answer is known . however, such passages are what is being sought when verifying a claim.
Approach: They propose a method that generates questions based on different focal points within a claim . they demonstrate that the method generates more relevant and informative questions .
Outcome: The proposed method outperforms previous work on a fact-checking question generation dataset on measurable evaluation metrics.
Tailoring Vaccine Messaging with Common-Ground Opinions (2024.findings-naacl)

Copied to clipboard

Challenge: Vaccine interventions aim to answer concerns expressed about vaccination.
Approach: They propose a dataset to evaluate how well responses are tailored to a common-ground opinion . they find that GPT-4-Turbo performs significantly better than others .
Outcome: The proposed dataset outperforms fine tuned LLMs on the task of tailoring vaccine responses to common-ground opinions.
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-modal Large Language Models have shown remarkable progress in visual contexts, yet their ability to convert visual figures into executable code remains underexplored.
Approach: They propose to use a set of visual coding metrics to assess MLLMs' visual . pass rate, text-match ratio, and GPT-4V rating judgement to assess the quality of generated code and rendered images.
Outcome: The proposed benchmark includes 132 high-quality matplotlib plots across six plot types, as well as 150 and 86 plots from Python’s and R’s plotly libraries respectively, totaling 368 plots.
Asking Clarification Questions in Knowledge-Based Question Answering (D19-1)

Copied to clipboard

Challenge: Existing clarification datasets with limited annotated examples do not address ambiguous phenomena.
Approach: They propose a dataset that allows users to ask clarification questions using open-domain examples.
Outcome: The proposed model achieves better performance than strong baselines and provides new challenges.
CodeExp: Explanatory Code Document Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing code-to-text generation models produce only high-level code summaries that do not capture implementation-level choices essential for these scenarios.
Approach: They propose a code explanation generation task that uses code docstrings to refine models.
Outcome: The proposed model can generate well-structured long docstrings comparable to human-written ones.
Image Caption Generation for News Articles (2020.coling-main)

Copied to clipboard

Challenge: Existing work on news-image captioning requires a joint understanding of image and text.
Approach: They propose a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption.
Outcome: The proposed model outperforms the state-of-the-art model and improves the quality of news-image captions.
Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for redacting offensive comments into non-offensive ones are inadequate to detect hateful content on social media platforms.
Approach: They propose a method for transforming offensive comments into non-offensive ones using a Retrieve, Generate and Edit unsupervised style transfer pipeline.
Outcome: The proposed method outperforms existing models on automatic metrics and human evaluations and consistently performs well on all automatic evaluation metrics.
Chain-of-Interactions: Multi-step Iterative ICL Framework for Abstractive Task-Oriented Dialogue Summarization of Conversational AI Interactions (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have introduced paradigm-shifting approaches in natural language processing, yet their transformative in-context learning (ICL) capabilities remain underutilized, especially in customer service dialogue summarization.
Approach: They propose a single-instance, multi-step framework that orchestrates information extraction, self-correction, and evaluation through sequential interactive generation chains.
Outcome: The proposed framework outperforms existing models and prompts in the customer service dialogue summarization domain.
Enhancing Topic-to-Essay Generation with External Commonsense Knowledge (P19-1)

Copied to clipboard

Challenge: Existing methods for topic-to-essay generation are insufficient for generating novel, diverse, and topic-consistent paragraph-level text with a set of topics.
Approach: They propose to integrate commonsense from external knowledge base into the generator through dynamic memory mechanism and adversarial training to further improve topic-consistency.
Outcome: The proposed task is more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation.
Assessing Dialogue Systems with Distribution Distances (2021.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics focus on turnlevel quality, which is not well suited for open-end dialogue tasks.
Approach: They propose to measure the performance of a dialogue system by computing the distributionwise distance between its generated conversations and real-world conversations.
Outcome: The proposed metrics correlate better with human judgments than existing metrics on dialogue systems.
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering (2025.naacl-long)

Copied to clipboard

Challenge: Generative AI has made rapid advances in multimodal understanding and code generation.
Approach: They construct a first real-world benchmark for multimodal large language models that directly convert visual designs into code implementations by manually curating 484 diverse real-life webpages as test cases.
Outcome: The proposed model can generate code implementations that directly render into the given reference webpages, given the screenshots as input.
Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style Transfer (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for unsupervised text style transfer lack parallel data and difficulties in content preservation.
Approach: They propose a neural approach to unsupervised text style transfer using non-parallel data.
Outcome: The proposed approach can be trained end-to-end on two widely-used public datasets.
Variational Hierarchical User-based Conversation Model (D19-1)

Copied to clipboard

Challenge: Recent approaches to conversation response generation model speakers and utterances together but are too tailored to the speakers.
Approach: They propose a new conversation model with a stochastic variable conditioned on the speakers and affects the context.
Outcome: The proposed model outperforms existing models in generating appropriate conversation responses.
REAM♯: An Enhancement Approach to Reference-based Evaluation Metrics for Open-domain Dialog Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for open-domain dialogue systems are limited by the diversity of possible outcomings.
Approach: They propose a method to augment a reference set to improve reliability . they propose BLEU to measure similarity between a predicted response and a small set of references .
Outcome: The proposed model improves the reliability of reference-based metrics with augmented reference sets.
Exploring Question-Specific Rewards for Generating Deep Questions (2020.coling-main)

Copied to clipboard

Challenge: Recent question generation approaches use the sequence-to-sequence framework to optimize the log likelihood of ground-truth questions using teacher forcing.
Approach: They propose to optimize for QG-specific objectives via reinforcement learning to improve question quality.
Outcome: The proposed model improves the fluency, relevance, and answerability of generated questions.
Reinforced Multi-task Approach for Multi-hop Question Generation (2020.coling-main)

Copied to clipboard

Challenge: Empirical evaluation shows our model to outperform the single-hop question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE and human evaluation metrics for quality and coverage of the generated questions.
Approach: They propose a question-aware reward function to maximize the utilization of supporting facts in the context.
Outcome: The proposed model outperforms single-hop neural question generation models on automatic evaluation metrics and human evaluation metrics for quality and coverage of the generated questions.
Factual Dialogue Summarization via Learning from Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary.
Approach: They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization.
Outcome: The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy.
NLG-Metricverse: An End-to-End Library for Evaluating Natural Language Generation (2022.coling-1)

Copied to clipboard

Challenge: Natural language generation models are a key component of deep learning, says aaron eliott . he says it is crucial to develop and apply better metrics for NLG evaluation .
Approach: a new open-source library for NLG evaluation is created to facilitate researchers to judge the effectiveness of their models. the framework provides a living collection of NLG metrics in a unified and easy-to-use environment.
Outcome: a new open-source library for NLG evaluation aims to improve performance of models . the framework provides tools to apply, analyze, compare, and visualize the metrics .
Improving Question Generation With to the Point Context (D19-1)

Copied to clipboard

Challenge: Existing sequence-to-sequence neural models may not be able to identify answer-relevant context words for question generation.
Approach: They propose to model the unstructured sentence and the structured answer-relevant relation for question generation by combining to the point context and unstructure.
Outcome: Experiments show that the proposed model improves on the unstructured sentence and the structured answer-relevant relation.
What is wrong with you?: Leveraging User Sentiment for Automatic Dialog Evaluation (2022.findings-acl)

Copied to clipboard

Challenge: Existing metrics for dialog evaluation are trained on human annotations, which is cumbersome to collect.
Approach: They propose to use user sentiment and other information as proxy to measure the quality of previous dialogs.
Outcome: The proposed model is comparable to models trained on human annotated data.
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to argument summarization rely on single-pass generation, offering limited support for factual correction or structural refinement.
Approach: They propose a large language diffusion framework that iteratively improves argument summarization by sufficiency-guided remasking and regeneration.
Outcome: Empirical results show that Arg-LLaDA surpasses state-of-the-art baselines in 7 out of 10 evaluation metrics.
Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to generate implausible stories using plots are unnatural and oversimplify the characteristics of implusible machine-generated stories.
Approach: They propose to generate a more comprehensive set of implausible stories using plots . plots are structured representations of controllable factors used to generate stories .
Outcome: The proposed model improves the quality of generated implausible stories using plots . it shows that the evaluation metrics trained on the generated data correlate better with human judgments compared to baselines.
Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for generating mathematical word problems are lacking in educational assessment.
Approach: They propose an end-to-end neural model to generate diverse mathematical word problems from commonsense knowledge graph and equations.
Outcome: The proposed model outperforms the SOTA models in terms of evaluation metrics and topic relevance.
Understanding the effects of word-level linguistic annotations in under-resourced neural machine translation (2020.coling-main)

Copied to clipboard

Challenge: Using word-level linguistic annotations in under-resourced neural machine translation is challenging for many languages.
Approach: They propose to use word-level linguistic annotations to label source-language (SL) or target-language words to improve translation performance.
Outcome: The proposed language annotations outperform part of speech and morphological description tags in the target language, while the morpho-syntactic description tags improve the grammaticality of the output.
Towards Automatic Evaluation for Image Transcreation (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for evaluating image transcreation have relied on human evaluation.
Approach: They propose a suite of automatic evaluation metrics inspired by machine translation metrics . they identify cultural relevance, semantic equivalence and visual similarity as critical dimensions of image transcreation .
Outcome: The proposed evaluation metrics agree with human ratings across 7 countries.
Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for document simplification address complex factors such as technical terminology, metaphors, and overall coherence.
Approach: They propose a multi-agent framework AgentSimp for document simplification based on large language models that simulates collaboration among agents through roles played by multiple agents.
Outcome: The proposed framework produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles.
Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems (2020.coling-main)

Copied to clipboard

Challenge: Existing evaluation metrics are not designed to cope with this flexibility.
Approach: They propose to group the qualities into three groups to obtain a single metric called USL-H.
Outcome: The proposed metric achieves good correlations with human judgment and maintains its configurability towards different aspects and metrics.
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models lack adequate evaluations and prompting strategies for explainability.
Approach: They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks.
Outcome: The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods.
On Temperature-Constrained Non-Deterministic Machine Translation: Potential and Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on the non-deterministic properties of language models, but these properties remain under-explored in machine translation.
Approach: They propose a method that evaluates MT systems and identifies temperature-constrained non-deterministic MT as a distinct phenomenon.
Outcome: The proposed framework provides higher-quality candidates than Deterministic MT under temperature constraints.
Logic-Consistency Text Generation from Semantic Parses (2021.findings-acl)

Copied to clipboard

Challenge: Text generation from semantic parses is challenging due to the complexity of the inner logic and the lack of automatic evaluation metrics for logic consistency.
Approach: They propose a framework for logic consistent text generation from semantic parses that employs iterative training procedures and quality control.
Outcome: The proposed framework enhances logic consistency and human evaluation on two benchmark datasets.
RepEval: Effective Text Evaluation with LLM Representation (2024.emnlp-main)

Copied to clipboard

Challenge: Traditional metrics for automatic text evaluation are tailored to specific tasks, while LLM-based evaluation metrics are costly.
Approach: They propose a metric that leverages projections of LLM representations for evaluation.
Outcome: The proposed metric exhibits higher correlation with human judgments than previous methods on 14 datasets.
LINKAGE: Listwise Ranking among Varied-Quality References for Non-Factoid QA Evaluation via LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Non-factoid (NF) question answering is challenging to evaluate due to diverse potential answers and no objective criterion.
Approach: They propose a listwise NFQA evaluation approach that uses Large Language Models to rank candidate answers in a descending list of reference answers sorted by descending quality.
Outcome: The proposed method has higher correlations with human annotations than standard methods.
Multi-hop Question Generation with Graph Convolutional Network (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on text-based QG focus on generating SQuAD-style questions.
Approach: They propose a multi-hop question generation model that does context encoding in multiple hops with Graph Convolutional Network and encoder fusion via an Encoder Reasoning Gate.
Outcome: Empirical results show that the proposed model generates fluent questions with high completeness and outperforms baselines on automatic evaluation metrics.
Towards a Better Metric for Evaluating Question Generation Systems (D18-1)

Copied to clipboard

Challenge: Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models .
Approach: They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function .
Outcome: The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available .
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for image captioning are primarily designed for short captions and are not suitable for long captions.
Approach: They propose an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework.
Outcome: The proposed metric outperforms existing metrics and achieves superhuman performance on LongCap-Arena.
Learning to Rank Visual Stories From Human Ranking Data (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on visual storytelling (VIST) use automated evaluation metrics for text generation.
Approach: They develop a Vrank metric that repurposes human evaluation results for automatic evaluation.
Outcome: The proposed model is more accurate than existing metrics and is generalizable to textual stories.
Phrase-Level Action Reinforcement Learning for Neural Dialog Response Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for dialog agent training lack a robust action space for entangled information, which can cause bias and deviate from natural human language.
Approach: They propose phrase-level action reinforcement learning which allows the model to alter the sentence structure and content with the sequential action selection.
Outcome: The proposed model achieves competitive results with state-of-the-art models on the MultiWOZ dataset, indicating that it is effective for solving task-oriented dialogs.
Evaluation Dataset for Zero Pronoun in Japanese to English Translation (2020.lrec-1)

Copied to clipboard

Challenge: In natural language, we often omit some words that are easily understandable from the context.
Approach: They propose to use a dataset to evaluate whether translation models can resolve zero pronoun problems in Japanese to English translations.
Outcome: The proposed model can resolve the zero pronoun problem in Japanese to English translations.
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic evaluation metrics for summarization are insensitive to factual inconsistencies.
Approach: They propose an automatic evaluation protocol that detects factual inconsistencies in a model-generated summary.
Outcome: QAGS has higher correlations with human judgments of factual consistency than other evaluation metrics.
Analysing Zero-Shot Readability-Controlled Sentence Simplification (2025.coling-main)

Copied to clipboard

Challenge: Text simplification (RCTS) models often depend on parallel corpora with readability annotations on both source and target sides.
Approach: They propose to use instruction-tuned large language models for zero-shot RCTS to reduce reliance on parallel corpora with readability annotations on both source and target sides.
Outcome: The proposed model can generate sentences with the desired readability, but the model's limitations and characteristics of the source sentences impede it.
Revisiting Grammatical Error Correction Evaluation and Beyond (2022.emnlp-main)

Copied to clipboard

Challenge: Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems.
Approach: They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system.
Outcome: The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task.
Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for abstractive text summarization do not provide explicit interdocument relationships among source documents.
Approach: They propose a model that uses sparse attention based on the conversational structure and a multi-task training objective that predicts metadata features.
Outcome: The proposed model outperforms baseline models in terms of evaluation metrics but struggle to handle conflicts in source documents.
Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought Method (2023.acl-long)

Copied to clipboard

Challenge: Experimental results show that automatic summarization generates concise summaries that contain key ideas of source documents.
Approach: They propose to use Element-aware test sets to annotate news-related reference summaries to focus on more fine-grained news elements objectively and comprehensively.
Outcome: The proposed method outperforms state-of-the-art fine-tuned PLMs and zero-shot LLMs by +4.33/+4.77 on the two datasets, respectively.
Composing Ci with Reinforced Non-autoregressive Text Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to compose Ci are limited in handling the constraints of tune patterns . authors propose a non-autoregressive approach to generate Ci using a synchronous process .
Approach: They propose to compose Ci using a non-autoregressive approach that takes into account rigid formats . they propose to apply reinforcement learning to the generation process with rigid constraints .
Outcome: The proposed method outperforms baselines and previous studies on a Ci dataset . it allows the model to perform synchronous generation while maintaining the format and content requirement.
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate.
Approach: They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Outcome: The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
EmailSum: Abstractive Email Thread Summarization (2021.acl-long)

Copied to clipboard

Challenge: Recent years have brought about interest in the task of summarizing conversation threads.
Approach: They develop an email thread summarization dataset that contains human-annotated short and long email threads over a wide variety of topics.
Outcome: The proposed dataset contains human-annotated short (30 words) and long (100 words) summaries of 2,549 email threads over a wide variety of topics.
Personalizing Dialogue Agents via Meta-Learning (P19-1)

Copied to clipboard

Challenge: Existing personalized dialogue models use human designed persona descriptions to improve dialogue consistency.
Approach: They propose to extend Model-Agnostic Meta-Learning (MAML) to personalized dialogue learning without using persona descriptions.
Outcome: The proposed model outperforms baseline models in terms of human-evaluated fluency and consistency on a persona-chat dataset.
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement.
Approach: They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction.
Outcome: The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics.
DISK: Domain-constrained Instance Sketch for Math Word Problem Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for generating MWP text from equations are inflexible and require pre-defined templates.
Approach: They propose a neural model which generates MWPs from equations by constructing a Quantity Cell Graph from the retrieved MWp instance and reasoning over it.
Outcome: The proposed model performs impressively on educational MWP set and on human evaluation metrics.
Data Manipulation: Towards Effective Instance Learning for Neural Dialogue Generation via Learning to Augment and Reweight (2020.acl-main)

Copied to clipboard

Challenge: Current state-of-the-art neural dialogue models learn from human conversations . however, due to the open-ended nature of human conversations, the quality of training data varies .
Approach: They propose a data manipulation framework to augment and highlight effective training samples . they also propose to increase its manipulation skills through gradient descent with validation samples a reshaping framework to proactively restructure the data distribution towards reliable samples is also proposed .
Outcome: The proposed framework improves the performance of open-domain neural dialogue models with respect to evaluation metrics and human judgments.
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output.
Approach: They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation.
Outcome: The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output.
Active Evaluation: Efficient NLG Evaluation with Few Pairwise Comparisons (2022.acl-long)

Copied to clipboard

Challenge: Recent studies show that evaluating NLG systems using pairwise comparisons is expensive as the number of human annotations grows linearly with k.
Approach: They propose a framework to efficiently identify the top-ranked system by actively choosing system pairs for comparison using dueling bandit algorithms.
Outcome: The proposed framework reduces human annotations by 80% on 13 NLG evaluation datasets spanning 5 tasks .
R4C: A Benchmark for Evaluating RC Systems to Get the Right Answer for the Right Reason (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have revealed that reading comprehension (RC) systems learn to exploit annotation artifacts and other biases in current datasets.
Approach: They propose a task that requires giving answers and derivations to evaluate RC systems' internal reasoning.
Outcome: The proposed framework annotates 4.6k questions with 3 reference derivations and shows that it is reliable and compares with existing benchmarks.
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in conversational AI have been substantial, but developing real-time tasks guidance systems remains a challenge.
Approach: They propose a data curation pipeline that synthesizes dialogues from annotated egocentric videos and a suite of automatic evaluation metrics that validated through extensive human studies.
Outcome: The proposed framework synthesizes dialogues from annotated egocentric videos and validates them through extensive human studies.
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects (2025.findings-acl)

Copied to clipboard

Challenge: In order to evaluate large language models (LLMs), it is important to collect benchmark datasets in order to assess their multilingual performance.
Approach: They extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages/dialects.
Outcome: The proposed dataset covers 55 languages and provides best-performing MT systems in all 55 languages.
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation (2025.coling-main)

Copied to clipboard

Challenge: Generative large language models generate a high-dimensional probability distribution over all tokens in their vocabulary.
Approach: They conduct extensive sensitivity analyses to determine how hyperparameter choices shape the outputs of generative large language models.
Outcome: The proposed methods influence the distribution of diversity and coherence metrics in human-written text, but the optimal configurations vary across models and tasks.
FarExStance: Explainable Stance Detection for Farsi (2025.coling-main)

Copied to clipboard

Challenge: FarExStance is a new dataset for explainable stance detection in Farsi . it contains extractive explanations as evidence for stance labels and claims .
Approach: They propose a dataset for explainable stance detection in Farsi with extractive explanations as evidence.
Outcome: The proposed model is the most accurate on stance detection, while the best explanation is from few-shot Claude-3.5-Sonnet.
Extended Parallel Corpus for Amharic-English Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Existing approaches to automate the complex task of translation are tedious and expensive.
Approach: They describe acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus.
Outcome: The proposed corpus outperforms statistical machine translation models by six to seven BLEU points . the results show that the subword models outperformed word-based models by three to four BLUE points compared with the word-base models .
Evaluating the Knowledge Dependency of Questions (2022.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for MCQ generation focus on the n-gram based similarity of the generated MCq to the gold sample and disregard their educational value.
Approach: They propose to use a human survey to measure the MCQ’s answerability given knowledge of the target fact.
Outcome: The proposed methods measure the MCQ’s answerability given knowledge of the target fact.
PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing image captioning metrics are vulnerable to lexical perturbations, but they are not robust to such perturbations.
Approach: They propose a perturbation-robust multilingual CLIPScore which is a reference-free image captioning metric for multiple languages.
Outcome: The proposed metric outperforms baseline metrics in capturing lexical noise of all various perturbation types in all five languages while maintaining a strong correlation with human judgments.
Data Sampling and (In)stability in Machine Translation Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: a recent data sampling method skews the annotated data toward shorter documents, not necessarily representative of the full test set.
Approach: They examine different approaches to human evaluation and ranking of machine translation systems at the conference on machine translation . they propose a method that uses available labour budget to sample data in a more representative manner .
Outcome: The proposed method improves representation of document lengths and produces stable rankings of translation quality.
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability.
Approach: They propose to evaluate the planning ability of large language models and multi-modal counterfactual vision language models (VLMs) using a multi-factual household activity simulator and a chatGPT task description to evaluate their reasoning ability.
Outcome: The proposed benchmark evaluates the planning ability of multi-modal and counterfactual vision language models on a household activity simulator and a chatGPT task description.
Teaching Language Models To Gather Information Proactively (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models are often defaulted to passive responses or narrow clarifications when faced with incomplete or under-specified prompts.
Approach: They propose a new task paradigm where LLMs must identify gaps in context and strategically elicit implicit user knowledge through targeted questions.
Outcome: The proposed framework outperforms o3-mini on evaluation metrics and human annotators favor clarification questions and final outlines.
Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and Challenge (2023.acl-long)

Copied to clipboard

Challenge: CR is the ability to understand and navigate the world using basic knowledge and understanding shared by most people.
Approach: They propose to incorporate pretrained knowledge into NMT models and use them as robust testbeds for investigating CR in NMT.
Outcome: The proposed method improves the training of NMT models with high CR abilities and provides accurate evaluation metrics.
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to evaluate open domain dialogues have a one-to-many problem . existing approaches lack commonsense reasoning biases and perform poorly in domain-specific scenarios.
Approach: They propose a framework that leverages both a small, specialised model and LLMs for the evaluation of open-domain dialogues.
Outcome: The proposed framework achieves state-of-the-art performance in both classification and evaluation tasks and exhibits better correlation with human judgements.
Re-Examining Summarization Evaluation across Multiple Quality Criteria (2023.findings-emnlp)

Copied to clipboard

Challenge: a number of automated evaluation metrics are evaluated by multiple quality criteria, such as relevance, consistency, fluency and coherence.
Approach: They propose a method that removes the confounding variable and detects unreliable correlations.
Outcome: The proposed method detects unreliable correlations between QCs and human scores . it is based on a multi-QC setup, but it fails to detect summary corruptions .
kNN-LM Does Not Improve Open-ended Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Interpolation-based retrieval-augmented language models (LMs) are a subtype of retrieval augmented language model that computes the probability of the next token by interpolating between the softmax distribution of the original LM and a token distribution formed by retrieving over an external datastore.
Approach: They propose to interpolate the predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.
Outcome: The proposed methods do not exhibit improvements in open-ended generation quality, as measured by automatic evaluation metrics and human evaluations.
Rethinking Efficient Multilingual Text Summarization Meta-Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: a limited number of human annotations are required to evaluate multilingual summarization evaluation metrics.
Approach: They propose a multilingual meta-evaluation framework that uses machine translation systems to transform a monolingual metaevaluations dataset into multilingual versions.
Outcome: The proposed framework outperforms classical text-matching-based metrics in non-English languages.
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for text summarisation have restrictive token limits, limiting their effectiveness.
Approach: They propose a human-annotated data set for evaluating automatic factuality metrics . they propose 'longDocFACTScore' framework which can be extended to any length document .
Outcome: The proposed framework outperforms state-of-the-art metrics in evaluating long document summarisation data sets.
Meta-Evaluation of Sentence Simplification Metrics (2024.lrec-main)

Copied to clipboard

Challenge: Automatic Text Simplification (ATS) is a major natural language processing task that aims to help people understand complex text.
Approach: They propose to use a human-annotated dataset to study automatic text simplification models to determine which metrics to use when evaluating new models.
Outcome: The proposed models reconstruct the text into a simpler format by deletion, substitution, addition or splitting, while preserving the original meaning and correct grammar.
Revisiting Context Choices for Context-aware Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Recent work has cast doubt on whether context-aware machine translation models learn useful signals from context or are improvements in automatic evaluation metrics just a side-effect.
Approach: They propose to use separate encoders for source sentence and context as multiple sources for one target sentence to train context-aware machine translation models.
Outcome: The proposed model improves translation quality even with empty lines as context, but the correct context improves it and random out-of-domain context degrades it.
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models generate human-like text, making them unreliable for biomedical relation extraction tasks.
Approach: They propose to use Large Language Models as judges to evaluate biomedical relation extraction . they propose structured output formatting for LLM-generated responses that helps LLMs improve their performance by 15%.
Outcome: The proposed method improves LLM-Judges' performance by 15% . it is cheaper and more efficient than human evaluation metrics, the authors say .
Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Current writing agents rely on predefined workflows and rigid thinking patterns to generate outlines before writing . authors propose a framework for long-form writing agents built on heterogeneous recursive planning .
Approach: They propose a general agent framework that achieves human-like adaptive writing . they propose recursive task decomposition and dynamic integration of task types .
Outcome: The proposed framework outperforms state-of-the-art approaches on both fiction and technical report generation.
CourtEval: A Courtroom-Based Multi-Agent Evaluation Framework (2025.findings-acl)

Copied to clipboard

Challenge: Existing automated evaluation metrics like ROUGE and BLEU show low correlation with human judgments.
Approach: They propose a multi-agent evaluation framework that integrates multiple agents . they use ROUGE and BLEU to evaluate natural language models .
Outcome: The proposed evaluation framework outperforms the current state-of-the-art methods in two meta-evaluation benchmarks.
One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness (2026.acl-long)

Copied to clipboard

Challenge: et al., 2010) show that hub embeddings are close to many unrelated examples in high-dimensional embeddable spaces . cross-modal encoders that project different modalities into a shared space are useful for cross-module applications .
Approach: They propose a method for identifying the hub embedding and its corresponding hub text . they use images to evaluate cross-modal encoders that project different modalities into a shared space .
Outcome: The proposed method can identify a single hub embedding and its corresponding hub text . it achieves comparable or higher similarity scores than human-written reference captions in many images .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations